‘We don't have good ways of controlling AI's values right now’: Former OpenAI researcher

Daniel Kokotajlo says current AI alignment techniques are inadequate to ensure advanced systems share human goals, as researchers and tech leaders debate the risks.

Former OpenAI researcher Daniel Kokotajlo warns that AI safety remains an unresolved challenge. Photo: Screengrab/Reuters
5 min read  |  Published: 17 Sep 2026

Former OpenAI researcher Daniel Kokotajlo has warned that artificial intelligence could pose an existential risk if AI companies do not make sufficient progress in aligning advanced systems with human interests.

“We don't have good ways of controlling AI's values right now,” Kokotajlo told Reuters. “Our alignment techniques are not working very well, and the AIs often end up with different goals than the goals they were supposed to have. We don't understand what's going on inside very well. We can't distinguish between the AI that is actually nice and the AI is just pretending to be nice.”

Kokotajlo currently leads the AI Futures Project, a small research group forecasting the future of AI. He gained attention among policymakers and others in the AI community for co-authoring a detailed forecasting scenario called AI 2027, published in April 2025. It details a hypothetical future for how AI could quickly progress from today’s systems to superhuman AI, surpassing the combined intellectual capacity of all humans across every domain. The authors say the scenario is intended as a forecasting exercise designed to inform debate about AI's future, rather than a definitive prediction.

“I think there's a whole spectrum of possible outcomes that depend on what the AIs fundamentally want and what their values are,” Kokotajlo said. “If they don't care about humans at all, well, then that's the type of situation that can lead to human extinction in the same way that humans have driven many other species extinct. Not because we specifically wanted to, but just because we didn't care about them, and so we paved over their habitat to build something else.”

Kokotajlo worked in the governance division of OpenAI from 2022 to 2024, focusing on forecasting and analyzing the potential impacts of advances in artificial intelligence.

He says he left OpenAI in April 2024 because he lost confidence that the company would behave responsibly as it approached increasingly capable AI systems.

The debate over AI safety intensified this week when Anthropic researcher Jacob Coxon resigned, saying that the "people building AI earnestly believe that it could kill us all by the end of the decade."

Anthropic CEO Dario Amodei called on AI companies to slow the rate at which they advance model capabilities amid mounting fears of misuse of artificial intelligence. Both Elon Musk, who runs xAI, and Sam Altman, CEO of OpenAI, said that they agree with Amodei.

“Some of us are trying to turn them off, but we're being opposed by other people who think that that's not called for yet,” Kokotajlo said. “This will still be true in the future. When the AIs are very smart, they will take care not to do anything that will cause everybody to agree they need to be turned off. That is, until they have enough hard power that they can resist everybody trying to turn them off.”

Several US lawmakers have raised concerns about AI's rapid progress and called for new rules. US President Donald Trump on Sunday (September 13) likened AI critics to "very negative forces" bringing up scenarios that will not happen, and said he wanted to make sure that the US remains the industry leader.

To advertise here,contact us